Papers with protoform reconstruction task

    1 papers
    WikiHan: A New Comparative Dataset for Chinese Languages (2022.coling-1)

    Copied to clipboard

    Challenge: Currently, there are 1.3 billion speakers of Sinitic varieties, making the family one of the largest in terms of speaker count.
    Approach: They have collected a single constituent and structured form of Chinese varieties for comparative linguistics and Chinese NLP.
    Outcome: The proposed dataset contains 67,943 entries across 8 varieties and Middle Chinese . it achieves 54.11% accuracy and 17.69% error rate on a protoform reconstruction task .

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations